Models and Harnesses Distinguished: Why the Harness Cannot Be Internalized

Dev Hub Sep 6, 2026

Model and Harness Distinguished: Why Internalization of the Harness Is Impossible

Baoyu rejects the claim that the model has internalized the harness. What is internalized, he holds, is strategic knowledge of tool use; the harness, as the execution layer, actually launches programs, governs permissions, and maintains state. The two are complements, not substitutes.

To the claim that Astra has outgrown Codex and the coding harness has been internalized, Baoyu offers a principled rebuttal. Training does encode strategy—when to call tools, how to split tasks, how to recover from failure—yet the model's output remains, at bottom, tokens. Command execution, context-window management, permission enforcement, sandbox isolation, and interruption recovery are discharged by the harness. The analogy is apt: driving skill, however deeply internalized, is useless without a car; one is left only to mime the wheel. Stronger models beget more complex toolchains and, in turn, greater execution demands on the harness. Model and harness, then, have never stood as rivals; they stand as mind and hands.

The past two days were spent testing Astra, amid a stream of striking demonstrations of Astra operating Blender. These, however, prove the reverse: Codex is not unequal to Astra; Astra is dependent on Codex as its harness.

Strip away the harness layer that Codex supplies, and no degree of internalization would enable Astra to power on a machine and run Blender modeling on its own.

What is internalized is knowledge.

Each training run does fix a great deal into the weights: when to invoke tools, how to decompose tasks, how to retreat after failure, which code style is superior. These are strategy and judgment, and as such are internalizable. Astra, accordingly, needs no lengthy prompts instructing it in Blender operations; it already knows.

The model itself, however, is capable of precisely one act: the emission of tokens in response to a prompt.

Its utterance—“I will draw a line here”—is mere text. The actual execution of the command, the return of results to the model, context-window management, the determination of which actions require user approval, the confinement of the model within a sandbox against accidental deletion, and the restoration of state after interruption: these are the work of the harness.

The assertion that the harness has been internalized thus conflates two distinct matters. The harness layer is not internalizable: it figures nowhere in the model's inputs or outputs. A tool cannot be trained into weights; only the manner of its use can be.

In short: technique of tool use is internalizable; the tool itself is not.

Consider the accomplished driver: however thoroughly driving experience has been internalized, remove the car and nothing remains but the pantomime of a steering wheel.

That models are growing stronger is not without consequence for the harness.

Formerly, the harness performed much of the planning, decomposition, and fallback on the model's behalf. With those faculties now native to the model, the harness may be made leaner—less concerned with how the model should think, more with how commands are executed.

Conversely, a more capable model reaches for more tools and more complex undertakings, and the resulting demands on the harness—execution capacity, permission governance, state management—only grow.

There is, therefore, no case of Codex having been outgrown by Astra: it is the model, through the harness, that commands the computer. Where Astra's capabilities exceed Codex's present design, the correct conclusion is that Codex requires upgrading—not that the harness layer may be discarded.

Model and harness are, and always have been, complementary; substitution was never at issue.

The model is the mind; the harness, the hands. However acute the mind, it reaches the physical world only through its hands.